Papers by Emerson Cabrera Paraiso
HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing reward models for explaining hate speech are optimized for broad notions of safety, but they assign lower scores to contextually rich explanations. |
| Approach: | They propose a reward model that integrates interpretable signals to better align reward scores with the needs of hate speech explanation. |
| Outcome: | The proposed model outperforms general-purpose baselines and improves pair-wise preference. |